high fidelity video prediction
High Fidelity Video Prediction with Large Stochastic Recurrent Neural Networks
Predicting future video frames is extremely challenging, as there are many factors of variation that make up the dynamics of how frames change through time. Previously proposed solutions require complex inductive biases inside network architectures with highly specialized computation, including segmentation masks, optical flow, and foreground and background separation. In this work, we question if such handcrafted architectures are necessary and instead propose a different approach: finding minimal inductive bias for video prediction while maximizing network capacity. We investigate this question by performing the first large-scale empirical study and demonstrate state-of-the-art performance by learning large models on three different datasets: one for modeling object interactions, one for modeling human motion, and one for modeling car driving.
Reviews: High Fidelity Video Prediction with Large Stochastic Recurrent Neural Networks
Quality: The paper is technically sound. Claims are supported by experimental results. The experimental study tests the proposed method on three standard datasets with correct methodologies and evaluations. Model capacity comparisons are covered in the supplementary material. I believe those comparisons are important, and it is better if authors can include them in the main paper.
Reviews: High Fidelity Video Prediction with Large Stochastic Recurrent Neural Networks
All reviewers agree that the paper offers a solid contribution on evaluating the capacity of large neural nets for video prediction tasks. They authors examine a variety of different settings, and provide interesting results. One suggestion for improvement is that the authors evaluate SOTA networks on more complex data sets, and cases where some scenes have higher uncertainty.
High Fidelity Video Prediction with Large Stochastic Recurrent Neural Networks
Predicting future video frames is extremely challenging, as there are many factors of variation that make up the dynamics of how frames change through time. Previously proposed solutions require complex inductive biases inside network architectures with highly specialized computation, including segmentation masks, optical flow, and foreground and background separation. In this work, we question if such handcrafted architectures are necessary and instead propose a different approach: finding minimal inductive bias for video prediction while maximizing network capacity. We investigate this question by performing the first large-scale empirical study and demonstrate state-of-the-art performance by learning large models on three different datasets: one for modeling object interactions, one for modeling human motion, and one for modeling car driving.
High Fidelity Video Prediction with Large Stochastic Recurrent Neural Networks
Predicting future video frames is extremely challenging, as there are many factors of variation that make up the dynamics of how frames change through time. Previously proposed solutions require complex inductive biases inside network architectures with highly specialized computation, including segmentation masks, optical flow, and foreground and background separation. In this work, we question if such handcrafted architectures are necessary and instead propose a different approach: finding minimal inductive bias for video prediction while maximizing network capacity. We investigate this question by performing the first large-scale empirical study and demonstrate state-of-the-art performance by learning large models on three different datasets: one for modeling object interactions, one for modeling human motion, and one for modeling car driving.
High Fidelity Video Prediction with Large Stochastic Recurrent Neural Networks
Villegas, Ruben, Pathak, Arkanath, Kannan, Harini, Erhan, Dumitru, Le, Quoc V., Lee, Honglak
Predicting future video frames is extremely challenging, as there are many factors of variation that make up the dynamics of how frames change through time. Previously proposed solutions require complex inductive biases inside network architectures with highly specialized computation, including segmentation masks, optical flow, and foreground and background separation. In this work, we question if such handcrafted architectures are necessary and instead propose a different approach: finding minimal inductive bias for video prediction while maximizing network capacity. We investigate this question by performing the first large-scale empirical study and demonstrate state-of-the-art performance by learning large models on three different datasets: one for modeling object interactions, one for modeling human motion, and one for modeling car driving. Papers published at the Neural Information Processing Systems Conference.